Papers with speech editing

3 papers
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing (2025.emnlp-main)

Copied to clipboard

Challenge: Autoregressive language model for multilingual speech editing and zero-shot text-to-speech synthesis is available in 11 languages.
Approach: They introduce an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot text-to-speech synthesis across 11 languages.
Outcome: The model generates high-quality, natural-sounding speech, even with limited per-language data . it shows robust performance in diverse linguistic settings, even in limited per language data compared to other models .
VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild (2024.acl-long)

Copied to clipboard

Challenge: VoiceCraft is a token-infilling neural codec language model for speech editing and zero-shot text-to-speech evaluation.
Approach: They introduce a token infilling neural codec language model that performs on speech editing and zero-shot text-to-speech tasks.
Outcome: The proposed model outperforms previous models on speech editing and zero-shot text-to-speech tasks.
FluentSpeech: Stutter-Oriented Automatic Speech Editing with Context-Aware Diffusion Models (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for speech editing still suffer from over-smoothing problem and lack of robustness due to stutter.
Approach: They propose a stutter-oriented automatic speech editing model that incorporates sutter information into the hidden sequence.
Outcome: The proposed model achieves state-of-the-art performance on a speech recording dataset . it can improve fluency of stuttering speech in terms of objective and subjective metrics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations